[NV] Add H200 DeepSeek-V4-Pro AgentX recipes / [NV] 添加 H200 DeepSeek-V4-Pro AgentX 配方 - #2364
Conversation
Add one 8xH200 aggregated TP8 (EP1/DP1) DeepSeek-V4-Pro FP8 Dynamo-SGLang AgentX recipe with EAGLE MTP and HiCache, sweeping concurrency [1,2,4,8,16] over a fixed serving topology. Uses the Marlin MoE backend and Dynamo header-based session affinity (X-Dynamo-Session-ID). - New recipe agg-h200-tp8-mtp-kvoffload.yaml + master key dsv4-fp8-h200-dynamo-sglang-agentic-agg (decode num-worker 0 so per-GPU accounting reflects the single aggregated TP8 worker). - benchmark_lib.sh: skip the legacy nvext conv-aware CLI routing when a recipe opts into the header path (AIPERF_HTTP_X_DYNAMO_SESSION_ID_FROM_CORRELATION_ID=true); the existing AIPERF_USE_DYNAMO_CONV_AWARE_ROUTING opt-out and default are unchanged. - launch_h200-dgxc-slurm.sh: dsv4 fp8 model-path routing, srt-slurm v1.0.10 overlay for the agentic recipe, on-demand SGLang/nginx squash imports, and AgentX dataset / HF caches mounted into the multi-node agentic path.
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
2 similar comments
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30310586014 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30310920956 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30311003032 |
2 similar comments
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30311003032 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30311003032 |
|
/stage-results |
|
@cquil11 staged run 30311003032: https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-07-27~r30311003032 This shared staging slot remains available until the next |
The SGLang image is already staged on the H200 cluster at /data/containers/*.sqsh, like the dsr1 multinode path which imports nothing. Drop the import_squash helper and its now-unused NGINX_IMAGE; the multinode path maps SQUASH_FILE into srtslurm.yaml directly.
中文:合并 origin/main 并解决 perf-changelog 冲突
|
Revoking the standing `/reuse-sweep-run` authorization on this PR (removing the bare command comment from 2026-07-28). A bare `/reuse-sweep-run` is standing rather than one-shot, and it has been silently swallowing sweeps here. The run at the current head, 30505397990, shows the gate emitting That matters because a real code change landed after the authorization: commit The last real evidence, 30311003032 (5/5 multi-node agentic), is at Re-authorize with an explicit run ID once a fresh sweep lands. |
|
@csahithi the AgentX/AIPerf harness has been updated, please merge origin/main into your branch and refresh your submission. Additional tuning may be necessary depending on the config. I apologize for any inconvenience. This is an automated message. |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30506436017 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30586537651 |
Bring PR #2364 onto current main, including AIPerf agentx-v1.0.1 at b7b16cf851885567988a643282266bce74e34437, while re-appending only this PR's changelog entry. 中文:合并 origin/main 并解决性能变更日志冲突。将 PR #2364 更新至当前 main,包含 AIPerf agentx-v1.0.1(b7b16cf851885567988a643282266bce74e34437),并仅在文件末尾重新追加本 PR 的变更日志条目。
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30938672335 |
|
/stage-results 30938672335 |
|
@cquil11 staged run 30938672335: https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-08-04~r30938672335 This run remains available across future @cquil11 已将运行 30938672335 发布到预发布环境:https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-08-04~r30938672335 后续的 |
…00 DSV4 AgentX 聚合引擎上启用 SGLang 指标 The aggregated worker never set enable-metrics, so no sglang:-prefixed series reached the AIPerf server-metrics export and the published trace charts had no backend data behind them. The merged GB300 sibling (agg-gb300-tp4-mtp-kvoffload.yaml) already sets it. Also add AIPERF_REQUIRED_SERVER_METRIC_PREFIX so a future gap fails the run loudly instead of publishing a partial artifact, matching the same sibling recipe. AIPERF_USE_DYNAMO_CONV_AWARE_ROUTING=0 is included for parity with that recipe; it is a no-op here because AIPERF_HTTP_X_DYNAMO_SESSION_ID_FROM_CORRELATION_ID=true already short-circuits the conv-aware routing branch in benchmark_lib.sh. Requires a re-sweep: the current results were produced without engine metrics.
# Conflicts: # perf-changelog.yaml # runners/launch_h200-dgxc-slurm.sh
…ntX 通道固定 srt-slurm v1.0.38 v1.0.10 only injected AIPERF_SERVER_METRICS_URLS when the runner was an AIPerfBenchmarkRunner. This recipe uses benchmark.type: custom, whose CustomBenchmarkRunner is not that subclass, so AIPerf was never told where to scrape and the trace artifacts carried no sglang: series even with enable-metrics set on the engine. v1.0.38 wires the logical SGLang worker leaders' /metrics URLs for custom benchmarks too. It is the same release the GB300 dsv4 dynamo-sglang AgentX lane already runs, so the model, framework, and benchmark type all match a path known to publish backend metrics.
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=31435999410 |
|
/reuse-sweep-run 31435999410 |
|
As a PR reviewer and CODEOWNER, I have reviewed this and have:
Additional detail section:Disclosure — I authored part of what I am signing off. The three most recent commits on this branch are mine, pushed as a maintainer fix rather than by the PR author: Validation evidence. Run 31435999410 ran on the exact PR head What this run fixed, and how it was verified. The previous sweep on this branch produced trace artifacts with no backend engine series, so the published curve rendered incompletely. Two independent causes, both addressed and both confirmed in the run log rather than assumed:
Speculative decoding and acceptance length. MTP via EAGLE with Evals — not applicable, left unchecked. The Single-node recipe publication — not applicable. The one recipe in this PR is a multi-node Model and scenario scope. MODELS.md lists DeepSeek-V4-Pro as active for Agentic coding, with the MTP-only deprecation recorded as not yet enacted. This submission is the agentic MTP arm on the upstream No engine or serving-stack patching. No Signed: |
✅✅✅ Verdict: PASS ✅✅✅✅ Check 0 (CODEOWNER): PASS — Note: the sign-off discloses that the signer authored the branch's three most recent maintainer-fix commits; flagged here for the core maintainer's awareness per that disclosure. |
|
/stage-results 31435999410 |
|
@cquil11 staged run 31435999410: https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-08-10~r31435999410 This run remains available across future @cquil11 已将运行 31435999410 发布到预发布环境:https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-08-10~r31435999410 后续的 |
Add one 8xH200 aggregated TP8 (EP1/DP1) DeepSeek-V4-Pro FP8 Dynamo-SGLang AgentX recipe with EAGLE MTP and HiCache. Sweep concurrency [1, 2, 4, 8, 16] over a fixed serving topology using the Marlin MoE backend and Dynamo header-based session affinity (
X-Dynamo-Session-ID).agg-h200-tp8-mtp-kvoffload.yamland master keydsv4-fp8-h200-dynamo-sglang-agentic-agg;decode.num-workerremains 0 so per-GPU accounting reflects the single aggregated TP8 worker.main, including AIPerfagentx-v1.0.1atb7b16cf851885567988a643282266bce74e34437.中文说明
新增一个基于 8×H200 的聚合式 TP8(EP1/DP1)DeepSeek-V4-Pro FP8 Dynamo-SGLang AgentX 配方,启用 EAGLE MTP 与 HiCache。在固定服务拓扑下扫描并发度
[1, 2, 4, 8, 16],采用 Marlin MoE 后端和基于X-Dynamo-Session-ID请求头的 Dynamo 会话亲和性。agg-h200-tp8-mtp-kvoffload.yaml和主配置键dsv4-fp8-h200-dynamo-sglang-agentic-agg;decode.num-worker保持为 0,使单个聚合式 TP8 worker 的单 GPU 统计准确。main,其中包含 AIPerfagentx-v1.0.1,提交为b7b16cf851885567988a643282266bce74e34437。